Goto

Collaborating Authors

 repeated active learning


Online Submodular Set Cover, Ranking, and Repeated Active Learning

Neural Information Processing Systems

Online Ranking: At each round, the learner produces an ordered list of items, then suffers loss or receives reward. In this paper, loss is: the number of items needed to achieve some coverage objective Example: The cost at each round is the number of pages the user needs to view to deduce the complete information they desire. Repeated Active Learning is an interesting special case where the list consists of questions to ask or tests to perform. Here a reasonable loss is the number of the tests we need to perform before we can make a accurate diagnosis. For these applications we propose a new online learning problem we call online submodular set cover.


Online Submodular Set Cover, Ranking, and Repeated Active Learning

Neural Information Processing Systems

We propose an online prediction version of submodular set cover with connections to ranking and repeated active learning. In each round, the learning algorithm chooses a sequence of items. The algorithm then receives a monotone submodular function and suffers loss equal to the cover time of the function: the number of items needed, when items are selected in order of the chosen sequence, to achieve a coverage constraint. We develop an online learning algorithm whose loss converges to approximately that of the best sequence in hindsight. Our proposed algorithm is readily extended to a setting where multiple functions are revealed at each round and to bandit and contextual bandit settings.


Online Submodular Set Cover, Ranking, and Repeated Active Learning

Neural Information Processing Systems

We propose an online prediction version of submodular set cover with connections to ranking and repeated active learning. In each round, the learning algorithm chooses a sequence of items. The algorithm then receives a monotone submodular function and suffers loss equal to the cover time of the function: the number of items needed, when items are selected in order of the chosen sequence, to achieve a coverage constraint. We develop an online learning algorithm whose loss converges to approximately that of the best sequence in hindsight. Our proposed algorithm is readily extended to a setting where multiple functions are revealed at each round and to bandit and contextual bandit settings.